Papers by Hellina Hailu Nigatu

6 papers
Evaluating Machine Translation Datasets for Low-Web Data Languages: A Gendered Lens (2026.findings-acl)

Copied to clipboard

Challenge: afan oromo, amharic, and tigrinya are low-resourced languages . they are used for training, benchmarks, news, health, and sports . afono o'mara: quantity does not guarantee quality of MT datasets .
Approach: They investigate the quality of machine translation datasets for three low-resourced languages . they found a large skew towards the male gender in the datasets .
Outcome: The results show that training data has large representation of political and religious text, but benchmark datasets focus on news, health, and sports.
A Case Against Implicit Standards: Homophone Normalization in Machine Translation for Languages that use the Ge’ez Script. (2025.emnlp-main)

Copied to clipboard

Challenge: Homophone normalization is a pre-processing step used in Amharic natural language processing (NLP) but it also results in models that are unable to process different forms of writing in a single language.
Approach: They propose a method where normalization is applied to model predictions instead of training data and a scheme where normalized data is preserved in training.
Outcome: The proposed model achieves an increase in BLEU score of up to 1.03 while preserving language features in training.
Cognate Detection for Historical Language Reconstruction of Proto-Sabean Languages: the Case of Ge’ez, Tigrinya, and Amharic (2025.coling-main)

Copied to clipboard

Challenge: As languages evolve, we risk losing ancestral languages.
Approach: They propose to use cognates to reconstruct proto-languages from cognates in child languages that have likely evolved from the same word in the proto-linguistics.
Outcome: The proposed method is based on automatic cognate detection and in-context learning with GPT-4o to generate the proto-language from the cognates and use Sequence-to-Sequence models.
The Zeno’s Paradox of ‘Low-Resource’ Languages (2024.emnlp-main)

Copied to clipboard

Challenge: 'low resource' languages are understudied by the NLP community, while 'high resource' is referred to as 'achieved', while high-resource languages are referred .
Approach: They qualitatively analyzed 150 papers from the ACL Anthology and popular speech-processing conferences that mention the keyword ‘low-resource.
Outcome: The proposed analysis reveals that several interacting axes contribute to ‘low-resourceness’ of a language and why that makes it difficult to track progress for each individual language.
mRAKL: Multilingual Retrieval-Augmented Knowledge Graph Construction for Low-Resourced Languages (2025.findings-acl)

Copied to clipboard

Challenge: Knowledge Graphs are structured multirelational graphs that store factual knowledge.
Approach: They introduce a Retrieval-Augmented Generation (mRAKL) based system to perform mKGC.
Outcome: The proposed approach improves over a no-context setting with an idealized retrieval system.
Viability of Machine Translation for Healthcare in Low-Resourced Languages (2025.emnlp-main)

Copied to clipboard

Challenge: MT errors are more pronounced in low-resourced languages where human translators are scarce and MT tools perform poorly.
Approach: They propose to use a publicly available machine translation system to analyze machine translation errors in healthcare domains.
Outcome: The proposed system reduces errors in two low-resourced languages for healthcare.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations